Papers with classification of
Hate Speech and Offensive Language Detection in Bengali (2022.aacl-main)
Copied to clipboard
| Challenge: | Existing research on hate speech detection in English does not cover low-resource languages like Bengali. |
| Approach: | They develop an annotated dataset of 10K Bengali posts consisting of 5K actual and 5K Romanized Bengali tweets. |
| Outcome: | The proposed model outperforms other models on training actual and romanized datasets by interpreting the semantic expressions better. |
Automatic Orality Identification in Historical Texts (2020.lrec-1)
Copied to clipboard
| Challenge: | a set of general linguistic features are used to identify conceptually-oral historical texts . linguists recognize that there is also a lot of variation within discourse modes . |
| Approach: | They propose to use general linguistic features to identify conceptually-oral historical texts . they find they are useful for determining conceptuality of historical data as for modern data . |
| Outcome: | The proposed features are used to identify conceptually-oral historical German texts . the features are useful in determining conceptuality of historical data as they are for modern data . |
Centering the Margins: Outlier-Based Identification of Harmed Populations in Toxicity Detection (2023.emnlp-main)
Copied to clipboard
| Challenge: | toxicity detection models focus on marginalized groups, but they obscure harms faced by intersectional subgroups. |
| Approach: | They use outlier detection to identify text about people with demographic attributes distant from the "norm" they find model performance is worse for demographic outliers than non-outliers . |
| Outcome: | The proposed model performance is worse for outliers than non-outliers, the authors say . their analysis also shows that outlier analysis can identify harms faced by intersectional groups . |